Papers with normalized version

2 papers
Benefits of Data Augmentation for NMT-based Text Normalization of User-Generated Content (D19-55)

Copied to clipboard

Challenge: Social media texts are considered important language resources for several NLP tasks, but their use of non-standard words makes it difficult to process and analyze UGC.
Approach: They propose to use a Neural Machine Translation approach to normalize lexical variants to their canonical forms to overcome performance drop in UGC.
Outcome: The proposed approach overcomes a data bottleneck in Dutch, a low-resource language.
Ranking Human and LLM Texts Using Locality Statistics (2026.findings-eacl)

Copied to clipboard

Challenge: The paper extends the Data Movement Distance (DMD) metric defined to measure the locality in computer memory to text by defining a new term designed to better characterize low-frequency tokens.
Approach: They propose to define a normalized version of the Data Movement Distance (nDMD) term is designed to better characterize low-frequency tokens.
Outcome: The proposed normalized version outperforms baselines and improves performance on the English subset of the M4 dataset and the GenAI detection shared task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations